Papers by Govardana Sachithanandam Ramachandran
[CASPI] Causal-aware Safe Policy Improvement for Task-oriented Dialogue (2022.acl-long)
Copied to clipboard
| Challenge: | Recent advances in off-policy reinforcement learning methods that use offline data as against a simulator have proven to be sample efficient. |
| Approach: | They propose a batch-RL framework for ToD policy learning: Causal-aware Safe Policy Improvement (CASPI) that uses a mechanism to learn fine-grained reward that captures intention behind human response and offers guarantee on dialogue policy’s performance against a baseline. |
| Outcome: | The proposed framework outperforms the current state of the art on an end-to-end dialogue task using a multiwoz2.0 dataset. |